Papers with objective metrics

11 papers
Controllable Neural Dialogue Summarization with Personal Named Entity Planning (2021.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that our proposed framework generates fluent and factually consistent summaries under various planning controls using both objective metrics and human evaluations.
Approach: They propose a controllable neural generation framework that can guide dialogue summarization with personal named entity planning.
Outcome: The proposed framework generates fluent and factually consistent summaries under various planning controls using objective metrics and human evaluations.
Tractable & Coherent Multi-Document Summarization: Discrete Optimization of Multiple Neural Modeling Streams via Integer Linear Programming (2022.emnlp-industry)

Copied to clipboard

Challenge: Multi-document summarization generates summary of corpus of documents consisting of related topics.
Approach: They propose a generic framework to jointly consider coherence and informativeness in multi-document summarization and offers provisions to replace individual components based on the domain of source text.
Outcome: The proposed framework consistently performs better than baselines for objective metrics and human evaluation.
Coherent and Concise Radiology Report Generation via Context Specific Image Representations and Orthogonal Sentence States (2021.naacl-industry)

Copied to clipboard

Challenge: Neural models for text generation are often designed in an end-to-end fashion, limiting their practical usability in downstream applications.
Approach: They propose a method to compute image representations specific to each sentential context and exploiting diverse sentence states to ensure topical continuity and content diversity of generated radiology reports.
Outcome: The proposed method outperforms baselines on objective metrics and human evaluations by 18% and 29% respectively in the evaluation for informativeness and content ordering respectively.
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)

Copied to clipboard

Challenge: Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages.
Approach: They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers.
Outcome: The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language.
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks assess tools in isolation, overlooking challenges such as functional overlap and cross-server orchestration, which can lead to overly optimistic evaluations.
Approach: They propose a five-level benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents within a hierarchical Model-Context Protocol (MCP) ecosystem.
Outcome: The proposed framework evaluates end-to-end tool orchestration by agents in hierarchical Model-Context Protocol (MCP) environments.
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Chart question answering (CQA) is a multimodal task for evaluating the reasoning capabilities of vision-language models.
Approach: They propose a chart question answering benchmark that incorporates multilingual contexts and supports open-domain textual outputs.
Outcome: The proposed framework outperforms the previous three common CQA paradigms: instruction-following, OCR-enhanced, and chain-of-thought.
From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text (2021.acl-long)

Copied to clipboard

Challenge: a computational model for code-switching text is lacking in the corpus of real text.
Approach: They propose a neural machine translation model to generate Hindi-English code-switched sentences using monolingual Hindi sentences.
Outcome: The proposed model reduces perplexity on a language modeling task and improves on linguistic inference tasks.
Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation (2024.lrec-main)

Copied to clipboard

Challenge: despite recent advances in speech synthesis, the focus of research has been on high-resource languages like English.
Approach: They propose a framework that incorporates modeling of syntactic and acoustic cues associated with pausing patterns.
Outcome: The proposed framework generates natural speech even for longer and intricate out-of-domain sentences, despite training on short audio clips.
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating (2025.acl-long)

Copied to clipboard

Challenge: Existing LLMs focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity.
Approach: They propose a multi-dimensional evaluation system and an optimized debating framework . they propose to use coT reasoning enhancement, web-based Retrieval Augmented Generation to optimize across various dimensions.
Outcome: The proposed framework outperforms baseline models in argument quality assessment and debate process simulation by 57%.
TunArTTS: Tunisian Arabic Text-To-Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Historically, TTS relied on classical methods that proved expensive in terms of data storage and often resulted in robotic-sounding output known as concatenative speech.
Approach: They propose to extract a mono-speaker speech corpus from an online dictionary and use it to develop end-to-end TTS systems for the Tunisian dialect.
Outcome: The proposed system is based on two approaches: training from scratch and transfer learning.
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment (2026.acl-long)

Copied to clipboard

Challenge: ImmersiveTTS model synthesizes intelligible speech and environmental audio from natural language descriptions.
Approach: They propose an environment-aware text-to-speech model that integrates natural speech with environmental audio . the model explicitly models cross-modal interactions through a dual-stream stage .
Outcome: Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations